Papers with online evaluation

7 papers
MUG: Interactive Multimodal Grounding on User Interfaces (2024.findings-eacl)

Copied to clipboard

Challenge: Prior studies modeled multimodal UI grounding in one round, but such an interaction is inherently iterative.
Approach: They propose a task where a user and an agent collaborate on an interface screen . they use a dataset of 77,820 sequences of human user-agent interaction on mobile interfaces .
Outcome: The proposed task improves the absolute task completion by 18% over the entire test set and 31% over the challenging split.
GEMv2: Multilingual NLG Benchmarking in a Single Line of Code (2022.emnlp-demos)

Copied to clipboard

Challenge: Evaluations in machine learning rarely use the latest metrics, datasets, or human evaluation in favor of remaining compatible with prior work.
Approach: They propose to use the Generation, Evaluation, and Metrics Benchmark to integrate new evaluation methods into existing evaluations.
Outcome: The proposed evaluation infrastructure bridges the gap between the advantages of leaderboards and in-depth and evolving evaluations by allowing model developers to benefit from each other's work.
Document-based Recommender System for Job Postings using Dense Representations (N18-3)

Copied to clipboard

Challenge: 45% of job posting traffic is driven by recommender systems for job postings . a large-scale job recommendation system is needed to detect similarity between job posting and item-to-item based recommendations.
Approach: They propose to use dense vector representations to enhance a large-scale job recommendation system and rank job advertisements regarding similarity.
Outcome: The proposed method increases the click-through rate on job recommendations by 8.0%.
Learning an Unreferenced Metric for Online Dialogue Evaluation (2020.acl-main)

Copied to clipboard

Challenge: Existing tools for dialogue evaluation do not generalize to unseen datasets and/or need a human-generated reference response during inference.
Approach: They propose an unreferenced automated dialogue evaluation metric that uses large pre-trained language models to extract latent representations of utterances and leverages the temporal transitions that exist between them.
Outcome: The proposed model achieves higher correlation with human annotations in an online setting, while not requiring true responses for comparison during inference.
Predicting Long-Term Citations from Short-Term Linguistic Influence (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to quantify linguistic influence in timestamped documents are not informative about extent to which a paper affected subsequent publications.
Approach: They propose to quantify linguistic influence in timestamped document collections by estimating a Hawkes process with a low-rank parameter matrix and identify lexical and semantic changes using contextual embeddings and word frequencies.
Outcome: The proposed method is based on an online evaluation with incremental temporal training/test splits, in comparison with a strong baseline that includes predictors for initial citation counts, topics, and lexical features.
Training-Free Test-Time Contrastive Learning for Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Existing training-free alternatives to training-based models are static or depend on external guidance.
Approach: They propose a training-free adaptation framework that enables a frozen LLM to improve online by distilling supervision from its own inference experiences.
Outcome: The proposed framework outperforms existing test-time adaptation methods under online evaluation.
RISK: A Framework for GUI Agents in E-commerce Risk Management (2026.acl-long)

Copied to clipboard

Challenge: RISK is a framework designed to automate multi-step web interactions in e-commerce risk management.
Approach: a new framework is designed to build and deploy GUI agents for e-commerce risk management . RISK-R1 provides a scalable, domain-specific solution for automating complex web interactions .
Outcome: RISK provides a scalable, domain-specific solution for automating complex web interactions in e-commerce risk management.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations